Papers with data-driven methods

15 papers
EVIDENCEMINER: Textual Evidence Discovery for Life Sciences (2020.acl-demos)

Copied to clipboard

Challenge: EVIDENCEMINER is a web-based system that allows users to query a natural language statement and retrieve textual evidence from a background corpora for life sciences.
Approach: They propose a web-based system that lets users query a natural language statement and automatically retrieves textual evidence from a background corpora for life sciences.
Outcome: EVIDENCEMINER is a web-based system that lets users query a natural language statement and automatically retrieves textual evidence from a background corpora for life sciences.
Towards Unified Representations of Knowledge Graph and Expert Rules for Machine Learning and Reasoning (2022.aacl-main)

Copied to clipboard

Challenge: Empirical study shows superiority of proposed method over time-tested knowledge-driven and data-driven methods.
Approach: They propose a cognitive knowledge graph that unifies expert rules and relational facts as the substrate of machine learning and reasoning models.
Outcome: Empirical results show the proposed method superior to time-tested methods . the proposed model can perform both learning and reasoning with labeled data .
How Do Large Language Models Perform in Dynamical System Modeling (2025.findings-naacl)

Copied to clipboard

Challenge: Recent data-driven methods often use graph neural networks (GNNs) to learn interactions between objects.
Approach: They propose prompting techniques for dynamical system modeling and evaluate their performance . they find that large language models demonstrate competitive performance without training .
Outcome: The proposed methods show competitive performance without training compared to state-of-the-art methods in dynamical system modeling.
XferBench: a Data-Driven Benchmark for Emergent Language (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods to teach models to "language" are full of bias, toxicity, and potential intellectual property violations.
Approach: They propose a benchmark for evaluating the overall quality of emergent languages using data-driven methods.
Outcome: The proposed benchmark is based on utterances from the emergent language and is validated using human, synthetic, and emergentic language baselines.
Social Norms-Grounded Machine Ethics in Complex Narrative Situation (2022.coling-1)

Copied to clipboard

Challenge: Recent studies focus on data-driven methods to judge the ethics of complex real-world narratives but face two major challenges: they cannot handle dilemma situations due to a lack of basic knowledge about social norms; and they focus on sparse situation-level judgment regardless of the social norm.
Approach: They propose to complement a complex situation with grounded social norms by a norm-supported ethical judgment model in line with neural module networks to alleviate dilemma situations and improve norm-level explainability.
Outcome: The proposed model improves state-of-the-art performance on two narrative judgment benchmarks.
A Workflow for HTR-Postprocessing, Labeling and Classifying Diachronic and Regional Variation in Pre-Modern Slavic Texts (2024.lrec-main)

Copied to clipboard

Challenge: a workflow for classifying diachronic and regional language variation in medieval texts is currently being developed . the workflow is generic or language-agnostic, but can be applied to other historical languages as well.
Approach: They propose a workflow for classifying diachronic and regional language variation in medieval texts . they use handwritten text recognition and manual transcription to obtain the data .
Outcome: The proposed workflow covers HTR-postprocessing, annotating and classifying medieval texts . it is accessible to humanists with limited experience in research data infrastructures, analysis or NLP .
Can Large Language Models Mine Interpretable Financial Factors More Effectively? A Neural-Symbolic Factor Mining Agent Model (2024.findings-acl)

Copied to clipboard

Challenge: Existing factor mining models are inefficient and inefficient, resulting in a significant challenge to extract interpretable factors.
Approach: They propose a model that integrates the strengths of both neural and symbolic models for factor mining.
Outcome: The proposed model surpasses the SOTA RankIC and RankICIR in predicting S&P 500 returns on real-world stock market data.
Linguistically-driven Framework for Computationally Efficient and Scalable Sign Recognition (L18-1)

Copied to clipboard

Challenge: a new general framework for sign recognition from monocular video is presented . the framework exploits state-of-the-art learning methods while incorporating features based on what we know about the linguistic composition of lexical signs.
Approach: They propose a general framework for sign recognition from monocular video . they exploit state-of-the-art learning methods while incorporating features from linguistic information .
Outcome: The proposed framework exploits state-of-the-art learning methods while incorporating features based on what we know about linguistic composition of lexical signs.
Syntax-Aware Opinion Role Labeling with Dependency Graph Convolutional Networks (2020.acl-main)

Copied to clipboard

Challenge: Opinion role labeling (ORL) is a fine-grained opinion analysis task . due to the scarcity of labeled data, ORL remains challenging for data-driven methods due to its complexity and complexity.
Approach: They propose to integrate syntactic knowledge into ORL models by comparing and integrating different representations and using dependency graph convolutional networks to encode parser information at different processing levels.
Outcome: The proposed model achieves 4.34 higher F1 score than the current state-of-the-art.
Identifying Aspects in Peer Reviews (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to peer review are limited in how they identify aspects . a growing volume of peer review submissions is straining the process .
Approach: They propose a data-driven schema for deriving aspects from peer reviews . they propose augmented peer reviews and show how it can be used for community-level review analysis.
Outcome: The proposed approach can be used to support peer review, but lacks formal definition of aspect . it also shows that the choice of aspects can impact downstream applications .
A Large-Scale Corpus for Conversation Disentanglement (P19-1)

Copied to clipboard

Challenge: a dataset of 77,563 messages manually annotated with reply-structure graphs disentangles conversations and defines internal conversation structure.
Approach: They use a dataset of 77,563 messages manually annotated with reply-structure graphs to disentangle conversations and define internal conversation structure.
Outcome: The new dataset is 16 times larger than all previous datasets combined and includes adjudication of annotation disagreements and context.
Understanding the Language of Political Agreement and Disagreement in Legislative Texts (2020.acl-main)

Copied to clipboard

Challenge: Despite the fact that state-level legislation is rarely discussed, it has a dramatic influence on the everyday life of residents of the respective states.
Approach: They propose a large-scale dataset linking state bills and legislator information, geographical information about their districts, and donations and donors’ information.
Outcome: The proposed model improves over strong text-based models by integrating the state-level text and the legislative context.
How Does the Experimental Setting Affect the Conclusions of Neural Encoding Models? (2022.lrec-1)

Copied to clipboard

Challenge: Recent studies have shown that neural encoding models explore brain language processing using naturalistic stimuli.
Approach: They propose a block-wise cross-validation training method and an adequate data size for increasing the performance of neural encoding models.
Outcome: The proposed training method and data size can significantly decrease the performance of neural encoding models in the temporal and frontal lobes.
A Large Collection of Model-generated Contradictory Responses for Consistency-aware Dialogue Systems (2024.findings-acl)

Copied to clipboard

Challenge: Recent large-scale neural response generation models (RGMs) have made significant progress but still struggle to generate semantically appropriate responses.
Approach: They build a large dataset of model-generated contradictions for the first time and analyze the results to gain valuable insights into their characteristics.
Outcome: The proposed dataset significantly improves the performance of data-driven contradiction suppression methods.
MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct (2025.findings-acl)

Copied to clipboard

Challenge: a new framework for image-text instruction data evolution improves MLLM performance . lack of high-quality instruction data remains a major bottleneck in ML modeling .
Approach: They propose a multimodal instruction data evolution framework that iteratively enhances data quality through fine-grained perception, cognitive reasoning, and interaction evolution.
Outcome: The proposed approach improves MLLM performance in nine vision-language tasks while using significantly less data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations